Animal Cognition
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match Animal Cognition's content profile, based on 23 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Wang, S.; Li, J. W. J.; Ma, C.; Poon, E. S. K.; Sin, S. Y. W.
Show abstract
Long-term memory has been extensively documented in many animals across a range of ecological contexts. Far less is understood, however, about how long memories persist after active modification or how conflicting memories are adaptively regulated to enable flexible behaviour. Parrots are widely recognized for their high intelligence and cognitive capability, yet the study of memory has focused on only a handful of large parrot species. Here, using a binary choice symbolic system, we conducted three experiments on 28 rosy-faced lovebirds (Agapornis roseicollis) to investigate (1) memory retention of initial associative learning at intervals of two weeks, half a year, and one year; (2) reversal learning, where the rewarded and unrewarded symbols were swapped; and (3) memory retention, following reversal learning at two-week and half-year intervals. We revealed that (a) initial associative learning memory was strikingly persistent, remaining detectable after nearly 600 days; (b) reversal learning performance was comparable to that of large parrot species; and (c) while the post-reversal forgetting curve mirrored the pattern of initial learning, it exhibited a shorter retention duration and a lower retention plateau. These findings suggest that forgetting may serve as an adaptive mechanism that regulates the retrievability of conflicting information, rather than a simple erasure of memory traces. When the environment changes, the retrievability of previously functional but now maladaptive memories is suppressed, but their long-term storage is preserved for potential future use under analogous circumstances. Such dynamic regulations prevent the repetitive overwriting of information in a fluctuating environment, thereby fostering greater behavioural flexibility across dynamic scenarios.
Serda, A.; Canteloup, C.; Meunier, H.
Show abstract
Inequity aversion, the sensitivity to inequitable outcomes or processes, has been widely studied in nonhuman primates since 2003. Yet, findings remain debated as some could be explained by alternative mechanisms such as frustration or loss aversion, so-called individual contrasts. Previous work on nonhuman animals has restricted the definition of inequity aversion to the social domain. Here, we propose distinguishing between two different forms: a socially driven form, which depends on comparison with another individual, and an individually based form, caused by a discrepancy between ones own effort and the expected outcome. Using a task-based methodology, we manipulated both effort and the amount of reward to test for the existence of those different forms and distinguish them from individual contrasts. We presented seven Tonkean macaques (Macaca tonkeana) and three brown capuchins (Sapajus apella) with low and high effort tasks associated with low and high reward quantities. To test for the presence of individually based inequity aversion, subjects were tested alone in an individual phase. In some of the trials, they were rewarded less than they deserved for the task. To test for socially based inequity aversion, two individuals performed the same task in a social phase but, in some trials, one received a higher-value reward than the other. We recorded the latency to engage in the task as an indicator of reluctance. In the individual phase, macaques were slower to engage with inequitably rewarded tasks, a pattern consistent with an individually based expectation of equity. By contrast, capuchins were faster in this context, suggesting that their responses were more likely driven by individual contrast effects. In a context of social inequity, both species slowed engagement in tasks with inequitable rewards, suggesting an aversion to socially based inequity. These results demonstrate that such a methodology provides a means to study non-social components of inequity aversion. Future studies, conducted on larger and more diverse samples, with rigorous motivational controls and explicit tests looking at the understanding of the link between task and reward, are needed to confirm whether nonhuman primates can display inequity aversion independently of social comparison.
Wang, S.; Hung, C. Y. C.; Poon, E. S. K.; Sin, S. Y. W.
Show abstract
Behaviour innovation plays a pivotal role in a species adaptability to dynamic environments. Investigating innovative behaviour and its underlying mechanism is therefore crucial for elucidating the development of cognitive flexibility across animals. Animal personality--which shapes how individuals perceive and engage with their surroundings--could offer insights into individual variation in this process. This study used a three-step foraging puzzle to evaluate the innovation capacity in 28 rosy-faced lovebirds (Agapornis roseicollis), specifically examining their capacity to recombine individually innovated component behaviours into integrated, more sophisticated techniques. We found that nearly half of the individuals spontaneously innovated multiple component behaviours to solve novel puzzles. Crucially, when challenged with a more sophisticated task, they recombined these behaviours into functionally dependent sequences without prior social demonstration. We further identified sex, persistence, and asocial learning capacity as key predictors of innovative problem-solving performance, with females, persistent individuals, and superior asocial learners excelled at problem-solving. Our findings demonstrate that behavioural innovation is not a static event, but a dynamic process--modulated by physical, cognitive, and personality variables--in which behaviours are flexibly transferred and recombined into increasingly complex forms to enable rapid individual adaptation.
Cantwell, A.; Masuda, B.; Greggor, A.
Show abstract
One challenge of reintroduction programs is ensuring that captive individuals are prepared for life in the wild, including having the ability to forage on wild resources. Foraging proficiency can be especially challenging to foster in species with a wide dietary breadth or reliance on ephemerally available resources. A relatively unexplored obstacle to release readiness may be an animals hesitancy to approach or consume unfamiliar foods, i.e. their neophobia and dietary conservatism. In this study, we examined the responses of a conservation breeding population of the extinct-in-the-wild alal[a] (Corvus hawaiiensis), a dietary generalist, to novel and familiar native and non-native fruits. Birds were similarly willing to contact both novel and familiar fruits, suggesting little aversion to the presence and appearance of novel food items. However, birds showed a clear preference for contacting non-native fruits over native fruits, regardless of familiarity. Additionally, while birds were willing to contact novel items, they were less likely to consume novel native fruits, suggesting that dietary conservatism may limit incorporation of these resources in the wild. Finally, juvenile birds were less likely to approach the testing setup and contact items, suggesting that neophobia may constrain their exploration of novel environments and resources. We discuss the implications of alal[a] preferences for non-native foods and reluctance to consume novel native items for pre-release preparations and the goal of wild species recovery.
Wild, S.; Uyehara, I. K.; Todd, L. M.; Ravara, T. A.; Sih, A.; Smith, J. E.
Show abstract
Human-influenced environments pose challenges but also provide wildlife with anthropogenic resources. Individuals vary widely in their ability to exploit such resources, often as a function of behavioural type. However, we lack a clear understanding of how variation in behavioural traits influences stages of resource exploitation required to use anthropogenic resources. Using fully automated foraging puzzles, we examined how boldness and sociability influenced three aspects of resource exploitation - discovery, problem-solving, and overall performance - in two wild populations of California ground squirrels (Otospermophilus beecheyi). Bolder individuals discovered the resource earlier, solved the task faster and achieved higher performance, indicating that boldness promotes efficient exploitation of anthropogenic resources across multiple stages in the process. Greater sociability and more opportunities to observe conspecifics solving the task led to faster problem-solving, consistent with evidence for observational learning. Squirrels in the recreational-use population - with regular exposure to humans and anthropogenic food - were faster to discover the resource than those in a trail-use population - where human exposure was transient and no anthropogenic food available. Problem-solving latency and performance were consistent between populations. Our findings highlight how individual variation in behavioural traits drives performance in novel ecological contexts, providing a mechanistic understanding of behavioural plasticity in human-influenced environments.
Alvi, M. H.; Nathaniel, R. P.; Rai, K.; Ma, L.
Show abstract
Adapting behaviour when reward contingencies change is a core function of cognitive control, but the underlying trial-by-trial computations are hard to observe when tasks cue each rule or allow only one switch per session. We developed the Feature-Rule Switching Task (FRST), in which common marmosets (Callithrix jacchus) repeatedly switch, without cues and under a fixed stimulus set, among the visual features that earn reward, inferring each switch from feedback alone. All four animals acquired the task within two days and sustained several switches per session across months of testing. A reinforcement-learning model with a learned weighting of stimulus dimensions best explained choices in every animal, outperforming complexity-matched perseveration controls. Learning-rate estimates fell within the human range, and in almost every session the animals weighted a stimulus dimension rather than individual features alone, with a dimensional commitment comparable in strength to that of humans. FRST thus provides a primate paradigm, amenable to laminar recording, for tracking the dynamics of cognitive flexibility.
Downie, I.; Szyszka, P.; Hall, N. J.; Edwards, T. L.
Show abstract
In turbulent environments, odorants from different sources arrive at different times, potentially providing cues for odor source segregation. In several invertebrate species, short differences in odorant onset enable freely moving animals to discriminate odorant mixtures. In vertebrates, however, studies of sensitivity to odorant onset asynchrony have been conducted under highly constrained sampling conditions, such as with odor delivery tightly coupled to respiration. In this study, we investigated whether domestic dogs could detect odorant onset asynchrony in odorant mixtures under conditions that preserve key features of natural odor sampling. Dogs performed a discrimination task in which odor stimuli were presented as ongoing pulse trains that began independently of animal behavior, avoiding artificial synchronization of odor delivery with sniff cycles. Dogs were trained to discriminate between mixtures of two odorants with synchronous onsets and mixtures with asynchronous onsets. Of the dogs trained, one was able to discriminate odorant onset asynchronies as short as 633 ms. Dogs also displayed sensitivity to auditory stimulus onset asynchrony, discriminating auditory asynchronies as short as 30 ms. These results provide the first demonstration of temporal sensitivity in canine olfaction and the first evidence that vertebrates can use odorant onset asynchrony under conditions that permit free odor sampling.
Roy, D. J.; Burton, T. J.; Balleine, B.
Show abstract
Considerable evidence suggests that the motivational control of instrumental action depends on incentive learning; i.e., on the opportunity to learn how the value of the consequences or outcome of an action, (e.g., a specific food) varies under different motivational conditions (e.g., under different degrees of hunger). The current study investigated whether learning the values of high-protein and high-carbohydrate rewards under different degrees of protein and carbohydrate appetite is also necessary for these nutrient-specific appetites to exert control over instrumental performance. Experiment 1 gave differing consummatory experience to whey protein and polycose carbohydrate outcomes under protein and carbohydrate appetite and found that, without the opportunity for incentive learning, the performance of actions earning these outcomes was insensitive to a shift in appetite. However, once the opportunity for incentive learning was provided, the rats increased their instrumental performance on a lever that earned the whey outcome relative to the polycose lever when protein hungry and on the polycose lever relative to the whey lever when carbohydrate hungry. Experiment 2 assessed how these nutrient-specific states exerted this control; whether, once learned, nutrient values were immediately controlled by nutrient appetite or whether this was based on conditional control acquired during experience with the outcomes under different nutrient appetites. We found that exposure to an outcome under a single nutrient-specific state was not sufficient to establish state-specific control. Instead, establishing the conditional control of outcome value required exposure to both the whey and polycose outcomes under both protein and carbohydrate appetites.
Mason, S. L.; Walsh, S. L.; King, S. L.; Ridley, A. R.
Show abstract
Syntax was long considered to distinguish human language from other vocal systems, with parallels in non-human animals historically limited to song. However, song lacks discrete meaning, which is a crucial pre-requisite of linguistic syntax. Over the last two decades evidence of combinatoriality in the discrete, semantic calls of an array of taxa has accumulated, providing the opportunity to investigate potentially closer parallels to language. However, most examples remain limited to small repertoires of simple two-call sequences, preventing evidence of complex internal structuring like that seen in human sentences. The recent discovery that several species produce extensive repertoires of much longer call sequences, has provided the opportunity to investigate the full extent of syntactic structure in non-human call systems. Here we demonstrate that Western Australian magpies (Gymnorhina tibicen dorsalis) use multi-level structured ordering rules within their semantic call sequences and that these ordering rules are learned during development. Specifically, we find that calls within sequences up to 15 calls long depend on the two calls given prior and that independently produced segments ( phonemes), calls, and sequences recombine into longer structures, indicating hierarchical organisation. This represents the first evidence of multi-level non-adjacent organisation and learned syntactic structure in a semantic non-human system.
Dahl, C. D.
Show abstract
Research on prosocial behaviour in nonhuman primates faces two related but distinct problems. First, empirical findings are heterogeneous: studies report instrumental helping, targeted tool transfer, food provisioning, consolation, and collaborative action, whereas others report little or no concern for partner welfare in low-cost prosocial-choice paradigms. These differences may reflect paradigm-specific variation in recipient need, action cost, actor benefit, joint payoff, solicitation, relationship context, competition, affordance clarity, and response bias rather than disagreement about one mechanism. Second, the interpretive vocabulary is underconstrained. Terms such as altruism, empathy, prosocial motivation, helping, and collaboration are often defined from behavioural outcomes and then treated as if they identified distinct latent mechanisms. This manuscript proposes a computational reframing of the problem. Prosocial behaviour is formalised as socially gated action selection. In the model, instrumental helping, low-cost prosocial choice, costly recipient benefit, strict altruism-like outcomes, need-sensitive empathy-like responding, and coordinated joint action are task and outcome profiles generated by different configurations of independently specified variables. These include need detection, goal inference, affordance recognition, action cost, actor benefit, relationship and tolerance effects, competition, request context, social feedback, joint payoff, action selection, and motor or apparatus bias. An illustrative simulation shows how low-cost instrumental helping, cost- or risk-suppressed recipient benefit, false-positive prosocial choice, and reliable-partner collaboration can arise from the same architecture. The framework identifies which task variables must be measured or manipulated before stronger claims about other-regarding motivation, empathy-like responding, or collaboration are justified.
Robert, T.; Flett, E.; Le Lay, H.; Nicolas, M.; Nityananda, V.
Show abstract
In vertebrates, top-down visual attention is a cognitive process where internal goals modulate the tuning of peripheral sensory systems. This leads to increased perceived contrast to both goal-relevant objects and areas of the visual field that are attended. Such a system would also be beneficial to bees, enabling them to detect and recognise the most profitable flowers in their environment. We tested whether bumblebees possess a top-down attentional system resembling that seen in vertebrates. We trained two groups of bees to collect rewards under high contrast targets. To potentially induce a difference in attention while searching for the targets, one group received a higher concentration of sucrose rewards compared to the other. During tests, the targets were presented with a series of lower contrasts to measure the contrast sensitivity curves of the bees induced by the different learnt reward levels. We predicted a stronger effect of any attention-like process on contrast sensitivity in the high reward group. We also repeated this experiment with the neonicotinoid pesticide imidacloprid dissolved in the sucrose rewards to test whether this affects bee attention. Across all test contrasts, higher rewards significantly increased bee accuracy when locating targets, lowered contrast thresholds and reduced the latency to make first choices. Imidacloprid reduced bee accuracy but did not influence first choice latency. These results suggest that learnt floral rewards can influence bee behavioural contrast sensitivity in a manner resembling vertebrate top-down attention and that imidacloprid may modulate this through effects on their nervous system.
Burnham, Y. L.; Fawcett, T. W.; Leaver, L. A.; Higginson, A. D.
Show abstract
Classic foraging models and discounting tasks may oversimplify the decision-makers environment, resulting in a discrepancy between predicted and observed behaviour. In delay discounting tasks, animals typically steeply devalue larger-later (LL) outcomes, choosing smaller-sooner (SS) rewards after short delays. This steep discounting appears to be irreconcilable with natural foraging behaviour, where animals frequently endure long delays when travelling to find food, handling tough items, and storing food for future use. This apparent mismatch in behaviour has led to questions regarding the ecological validity of laboratory discounting tasks. Here, we developed a rich dynamic optimisation model to identify the conditions under which animals should choose LL or SS outcomes. In our model, a food-storing animal encounters food items differing in energy content and handling time and must decide which items to eat immediately and which to cache, so that it has enough stored food to survive winter. We simulated a range of environments, including laboratory conditions where foragers face negligible predation risk when searching for food and have a high probability of finding food, compared to more natural conditions where searching for food is risky, and food is harder to find. In line with preference reversals seen in the discounting literature, our model predicts that LL items should be rejected more often when the handling time for such items is increased, whereas SS items should be rejected more often when the handling time for both items is increased. Importantly, our model only predicts rejection of food items under parameter values that reflect laboratory conditions, supporting the notion that the steep devaluation of rewards seen in animals may be driven by the artificiality of traditional discounting tasks.
Yasueda, M.; Taira, M.; Akam, T.; Walton, M. E.; Doya, K.
Show abstract
Reinforcement learning theory formulates distinct decision-making strategies, including reactive model-free and deliberative model-based strategies. This study investigates how mice adjust their reinforcement learning strategies while learning decision-making in dynamic environments. Unlike previous studies that focused on behaviors after extensive training periods, we analyzed changes in learning strategies in the course of training of a two-step decision-making task with probabilistic state transition and fluctuating reward probabilities. Our statistical behavioral analysis showed that the stay-probability following common and rare transitions diverged with training, a signature of strategies that utilize knowledge of task structure. We fit various reinforcement learning strategies to behavioral data and found that structure-informed strategies became increasingly dominant in their behaviors during training. Whereas previous studies emphasized transition from goal-directed to habitual strategies after extensive training, which were often associated with model-based and model-free strategies, respectively, our results newly demonstrate a shift from model-free to structure-informed strategies in early training in mice. Author summaryReinforcement learning theory allows us to examine how we make decisions and what approaches we use to optimize rewards. Most previous research, however, has examined animal behavior only after extensive training. Here we analyzed how mice adjust their reinforcement learning strategies as they are trained in a two-step decision-making task. Initially, mice relied on reactive model-free strategies, but as training progressed, their behavior began to incorporate knowledge of task structure. While previous studies suggested transition from model-based to model-free strategies with extensive training, our study revealed the opposite in the early stage of training.
White, T. E.; Rojas, B.; Mappes, J.; Kemp, D. J.
Show abstract
Natural scenes are visually complex, with dynamic variation in colour and luminance presenting challenges for detecting salient objects. However, the cues guiding human detection under such natural conditions remain largely untested. We placed custom-built coloured targets in a forest-like setting and quantified their contrasts against natural backgrounds using human visual models. Across human observers (n = 51), detection performance rose sharply with hue and saturation contrast, whereas luminance contrast--long considered a primary driver of visual search--contributed little explanatory power. Strikingly, even large changes in ambient light chromaticity barely altered detection, despite the patchy illumination of the habitat. Under these naturalistic conditions, hue and saturation contrasts were stronger predictors of detection than luminance contrast, providing a rare, field-based test of visual models widely applied across human and animal vision research.
Rossi, N.; Nicholls, E.
Show abstract
Environmental predictability influences the value of information acquired through experience, yet relatively little is known about how instability in resource characteristics influences behavioural organisation during foraging. We tested whether repeated changes in floral orientation, a manipulation of environmental predictability, affect pollen foraging in bumblebees (Bombus terrestris) by exposing naive workers to either stable floral conditions (single flower orientation) or repeated inter-trial changes in flower orientation (three alternating flower orientations), under constant resource availability. We quantified pollen collection, foraging efficiency, revisitation behaviour, floral coverage and sonication behaviour across three successive foraging trials of either constant or variable flower orientation, before assessing performance in a common post-treatment preference test in which bees were offered all three flower orientations and higher pollen rewards. Environmental instability altered the organisation of foraging behaviour. Bees exposed to unstable floral conditions progressively reduced flower revisitation behaviour and were less likely to perform sonication, although floral coverage, defined as the number of unique flowers visited, remained unchanged. Contrary to our predictions, instability had only weak immediate effects on pollen acquisition and foraging efficiency compared to bees tested under stable conditions. However, previous exposure to instability generated carry-over effects in the common post-treatment preference test. Bees previously exposed to unstable conditions were significantly less likely to return with pollen and consequently collected less pollen overall than bees previously exposed to stable conditions. Our results demonstrate that environmental instability can influence pollen foraging in ways that are not captured by immediate measures of performance. Although behavioural adjustments appeared to buffer short-term consequences during repeated foraging trials, carry-over effects were evident when bees were later tested in a common high-reward, multi-orientation floral array. These findings highlight the importance of considering both behavioural flexibility and carry-over effects when evaluating how organisms respond to changing environments.
Jiwa, M.; Myles, D.; Bennett, D.
Show abstract
Recent research has suggested that the availability of non-instrumental information about the outcome of a risky choice increases risk appetite. In this study, we aimed to perform a conceptual replication of these findings and to examine the cognitive mechanisms underlying this effect. Across two experiments (N = 150, 102), we presented participants with mathematically fair gambles and allowed them to choose the size of their bet. Between trials, we varied the presence of non-instrumental information that would reveal the outcome ahead of time. In both experiments, we did not find consistent evidence for an effect of the availability of non-instrumental information on bet size. These findings suggest that the previously reported effects of non-instrumental information on risk appetite may have been an idiosyncratic feature of experimental design, rather than a more general phenomenon that characterises human decision making under risk.
Mason, S. L.; Ridley, A. R.
Show abstract
Growing evidence of animals combining discrete, meaningful calls into sequences--a feature once thought unique to linguistic syntax--has presented the opportunity to investigate the evolutionary origins of syntactic communication. The arbitrary assignment of meaning to words marks an important step in human language evolution, and a necessary precursor to generating further meaning through sentences. Studying how other animals that produce call sequences learn the meaning of these signals could help shed light on how referentiality and semantic combinatoriality evolved. Given the presence of meaningful call sequences has only recently been revealed in several non-human animals, ontogenetic studies of the comprehension of these vocalisations are, to date, non-existent. Western Australian magpies (Gymnorhina tibicen dorsalis) combine discrete calls into a diverse array of call sequences. Recent evidence shows these sequences are socially learned, but the developmental stage at which fledglings respond correctly to them remains unstudied. We performed playbacks of a discrete alarm call and call sequence to fledglings over the course of their first 18 weeks out of the nest, identifying when they differentiate between the low-level disturbance associated with the discrete call and the high-grade aerial threat associated with the sequence. Fledglings showed immediate vigilance to both vocalisations but exhibited significantly greater vigilance and upward scanning following the sequence. Critically, fledglings showed this response to the sequence from the first week of testing, with no effect of age on the response to either vocalisation. These findings suggest that comprehension precedes production of sequences in magpies and that sequence meanings are either learned rapidly or have an innate basis. While further investigation is essential, this study offers the first empirical insight into the ontogenetic emergence of combinatorial comprehension in a non-human animal.
Kittur, M.; Zhang, A.; Bryce, N.; Yousif, S.
Show abstract
Human spatial representations are often assumed to represent Euclidean properties such as length, distance, and angle. Here we test an alternative (but not mutually exclusive) possibility - that spatial memory is structured primarily around topological relations. Across four experiments, adults and children memorized simple letter-like figures and reproduced them by drawing, allowing the contents of their spatial representations to be revealed directly. Drawings showed systematic distortions of metric features, including strong biases of angles toward 90{degrees} and compression of line length towards an average value. In contrast, topologically critical features -- such as T-junctions and holes -- were reliably preserved, even relative to closely matched but topologically irrelevant features like L-junctions. These effects were magnified in a serial reproduction paradigm, in which participants iteratively generated new drawings from previous participant drawings: At the end of each mnemonic chain, figures converged on simplified topological structures as metric detail degraded. Similar patterns were observed in children aged five to eight years. Together, these findings suggest that basic topological relations may function as primitive building blocks of human spatial representation, with metric detail encoded secondarily. Significance statementThe iconic map of the London Underground is one of the most famous maps in history, yet something special about it goes unnoticed: it is not a veridical representation of space. Distances are arbitrary, and angles are presented only in coarse terms. Yet the ubiquity and appeal of such maps suggests that topological representation is intuitive -- as if the mind is keen to receive information in exactly this way. Here, using drawing as a tool, we show directly that the most primitive form of spatial representation appears to be a topological skeleton. Remarkably, even children as young as five represent spatial structure in topological terms, with roughly the same fidelity as adults -- pointing to an underappreciated building block of spatial representation.
Dai, Y.; Seielstad, A. K. L.; Hinman, J. R.
Show abstract
Social isolation has profound effects on behavior and cognition, but these effects can differ across sex and behavioral domains. We asked how social isolation shapes foraging decisions in male and female rats using a patch-leaving task that combines spatial navigation with sequential stay-or-go choices across varying travel costs and reward depletion rates. All groups scaled patch residence times with travel cost in a manner consistent with the marginal value theorem yet consistently overstayed beyond the optimal leaving time. The magnitude of overstaying was shaped by a strong interaction between sex and social status, with socially isolated females leaving patches closest to optimal and consuming food at the highest rate. Their foraging was accompanied by a coherent spatial and behavioral profile where the socially isolated females spent more time idle, preferentially occupying protected regions near the door and corridor, and avoiding the exposed patch center during both active foraging and periods of idling. These patterns are consistent with a conservative, safety-oriented strategy that simultaneously minimizes exposure and maximizes caloric return. Social isolation does not uniformly impair cognition but can selectively bias female rats toward efficiency-maximizing foraging decisions, consistent with the ecological pressures faced by outcast females in wild rat colonies.
Claeys, W.; Ruuskanen, V.; Mathot, S.
Show abstract
When we feel restless and easily distracted, continuously switching tasks (exploration), our pupils tend to be large. In contrast, when we are calmly focused on a single task (exploitation), our pupils tend to be small. According to the Adaptive Gain Theory (AGT), a switch from exploitation to exploration is associated with an increase in norepinephrine in the locus coeruleus, which in turn triggers pupil dilation. However, the AGT does not provide a functional explanation of why exploration triggers pupil dilation. One possibility is that visual sensitivity, which increases with pupil size, is especially important during exploration. We set out to provide evidence consistent with this functional explanation, as well as to replicate two key previous results. Participants performed a four-armed bandit task, which induces both exploration and exploitation behavior. During the task, participants also needed to detect an occasional and unpredictable near-threshold peripheral flash. We replicated two key results: pupils were larger during exploration than during exploitation; and increased pupil size (overall, independent of exploration status) was associated with increased visual sensitivity. However, most importantly, we did not find that visual sensitivity was higher during exploration than during exploitation; probably, the reliable-yet-tiny increase in pupil size during exploration was too small to affect visual sensitivity. We conclude that key previous results are replicable; however, common experimental paradigms, such as the four-armed bandit task, induce only small changes in exploration behavior. Therefore, more powerful paradigms are required in order to test functional explanations of pupil-size changes during exploration and exploitation.